Conversation
593f473 to
1d542c7
Compare
…cipe
Replace the hand-rolled bash server script
(dsv4_fp4_b200_vllm_cache_sources_mtp.sh) with a native srt-slurm
recipe at dsv4/vllm/b200-fp4-mtp/cache-sources.yaml. Three variants
select by CONC/KV_OFFLOADING: TP8 c8 native DRAM, TP8 c14 Simple NVMe,
TP8 c14 native DRAM+NVMe. NVMe variants use host_setup to
create/teardown /scratch/inferencex-kv-$SLURM_JOB_ID and
container_mounts with {job_id} templating.
Each search-space row in nvidia-master.yaml now carries srt-recipe:,
routing to the native-single-node launch path in the b200-nscale
launcher. The bash-specific routing, /ix mount override, NVMe directory
management, and GPU_MEMORY_UTILIZATION=0.85 export are removed from
launch_b200-nscale-slurm.sh (gpu-memory-utilization is now 0.85 in the
recipe).
A new srt-slurm patch (pr56318-overlay-validation.patch) adds
pr56318-overlay-check.sh to configs/patches/, referenced by the
recipe's setup_script to verify overlay checksums before engine start.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36900160302 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36900160302 |
|
/use 36272389390 |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Rerun vLLM #56318 at
6ce00ee2ebcf29076d671786c36bf8f8ca71bd0bon DeepSeek-V4-Pro.Previous revision: all four points passed. This rerun uses the updated attribution code.